Papers with visually grounded reasoning
Cultural Concept Adaptation on Multimodal Reasoning (2023.emnlp-main)
Copied to clipboard
| Challenge: | Past methods focused on multilingual and multimodal capabilities, and the improvement of multicultural competence is still an unexplored problem. |
| Approach: | They propose an annotation-free method for cultural-concept adaptation and construct a concept mapping set to facilitate model's comprehension of cultural-consensual mappings. |
| Outcome: | The proposed method outperforms baseline models on zero-shot and few-shot settings on five languages and cultures. |
Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms (2023.emnlp-main)
Copied to clipboard
Seungju Han, Junhyeok Kim, Jack Hessel, Liwei Jiang, Jiwan Chung, Yejin Son, Yejin Choi, Youngjae Yu
| Challenge: | NormLens is a visual-grounded framework for understanding commonsense norms . state-of-the-art models are not well-aligned with human annotation, we show . |
| Approach: | They propose a visual-grounded framework to study commonsense norms by NormLens . they find that models are not well-aligned with human annotation . |
| Outcome: | The proposed model judgments and explanations are not well-aligned with human annotations. |
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts. |
| Approach: | They propose a novel agentic framework that explicitly performs visual reasoning directly within the chart’s spatial domain. |
| Outcome: | The proposed framework achieves state-of-the-art accuracy on the ChartBench and ChartX benchmarks surpassing prior methods by up to 16.07% absolute gain overall and 17.31% on numerically intensive queries. |